Papers with inference speed-up approach

    1 papers
    RefreshKV: Updating Small KV Cache During Long-form Generation (2025.acl-long)

    Copied to clipboard

    Challenge: Existing methods for generating long sequences of tokens are expensive and require memory and computation resources.
    Approach: They propose a method that alternates between full context attention and attention over a subset of input tokens during generation.
    Outcome: The proposed method achieves comparable speedup to eviction-based methods while improving performance for various long-form generation tasks.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations